Papers with Chinese contexts
BotChat: Evaluating LLMs’ Capabilities of Having Multi-Turn Dialogues (2024.findings-naacl)
Copied to clipboard
Haodong Duan, Jueqi Wei, Chonghua Wang, Hongwei Liu, Yixiao Fang, Songyang Zhang, Dahua Lin, Kai Chen
| Challenge: | Modern Large Language Models (LLMs) facilitate high-quality, multi-turn dialogues with humans, but human-based evaluation of such a capability requires substantial manual effort. |
| Approach: | They propose to evaluate LLMs' ability to emulate human-like, multi-turn conversations using an LLM-centric approach. |
| Outcome: | The proposed model emulates human-like, multi-turn conversations using an LLM-centric approach. |
Error-Robust Retrieval for Chinese Spelling Check (2024.lrec-main)
Copied to clipboard
| Challenge: | Chinese Spelling Check (CSC) aims to detect and correct spelling errors in Chinese texts . current methods may not fully leverage existing datasets, resulting in insufficient annotated data . |
| Approach: | They propose a plug-and-play retrieval method with error-robust information for Chinese Spelling Check . they employ multimodal representations that fuse phonetic, morphologic, and contextual information . |
| Outcome: | The proposed method improves on the SIGHAN benchmarks on Chinese spelling check (CSC) the proposed method is based on training data and lacks adequate parallel corpora . |
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing large language models (LLMs) are prone to misuse and misinformation, posing serious compliance risks. |
| Approach: | They propose a bilingual red-teaming benchmark to test an LLM’s refusal of requests that violate financial compliance. |
| Outcome: | The proposed benchmark is based on real-world financial crime cases and ethical violations and includes 14 subcategories covering financial crimes and ethical breaches. |
Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch (2025.acl-long)
Copied to clipboard
Xueru Wen, Jie Lou, Zichao Li, Yaojie Lu, XingYu XingYu, Yuqiu Ji, Guohai Xu, Hongyu Lin, Ben He, Xianpei Han, Le Sun, Debing Zhang
| Challenge: | Existing Chinese resources are small in scale and limited to specific domains, making them insufficient for LLM post-training. |
| Approach: | They propose a Chinese-annotated reward model and a preference dataset to address this gap . they evaluate Chinese RMs on CheemsBench and construct an RM that captures human preferences . |
| Outcome: | The proposed RM achieves state-of-the-art performance on CheemsBench and CheeMePreference. |
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset (2026.acl-long)
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have expanded capabilities beyond text understanding . a novel Chinese financial multimodal evaluation benchmark is used to evaluate LVLM capabilities . |
| Approach: | They propose a Chinese financial multimodal evaluation benchmark to evaluate LVLMs' capabilities . the model has an overall accuracy of 66.11% and an average score of 77.18 . |
| Outcome: | The proposed model achieves an overall accuracy of 66.11% on the question answering task and an average score of 77.18 on detection, recognition, and information extraction tasks. |